Repository navigation
Refactor out inlined values in builtin materialized views - #39468
SangJunBak wants to merge 7 commits into
Conversation
3db1f61 to
03701c4
Compare
03701c4 to
631ddf2
Compare
0ce5fa8 to
a3430b2
Compare
…exes - We need to extract this into its own view such that later, we can extend and reuse mz_builtin_indexes without changing the schema of mz_indexes since it's an mz_catalog relation and stable - Initially we thought we had to inline the builtin VALUES into the materialized view, otherwise a silent change to an upstream view (i.e. mz_builtin_indexes) would silently create an error in the materialized view. However we've since realized that as long as the schema of the upstream view stays the same, any updates to it will simply sink into the materialized view. If the schema were to change however, e.g. we ASSERT NOT NULL on the mv but have the upstream view provide a NULL value, then the mv would error when queried. However, we have tests that actually query each builtin so this regression would be caught.
…indexes - Because we don't need to inline VALUES into the DDLs of our builtin materialized views, we can create a new mz_builtin_log_indexes such that both mz_sources and mz_indexes can query it rather than get sent an iterator of the builtin logs. - We simply extract inline VALUES code into a new builtin view called mz_builtin_logs then reference it in mz_builtin_indexes - This comes with the nice benefit of not having to dynamically generate each mz_indexes and define the order statically. - We can also turn make_mz_indexes into a static MZ_INDEXES like the rest of our builtins - In a later commit, we'll making the same refactor for mz_sources
- In mz_sources, rather than inline the builtin sources and logs as a VALUES list, read them from mz_builtin_sources - mz_builtin_sources already existed inside "make_builtin_sources", however it was just never used and was dead code. This is why we delete code in make_mz_sources but don't see a positive diff in this commit. However if you were to look inside make_builtin_sources, the same code we deleted exists there.
…tin_log_indexes - Move `privileges` to `mz_builtin_log_indexes` from `mz_builtin_sources` - Instead of embedding the builtin logs to mz_sources, just join on `mz_builtin_log_indexes`
Simply move make_mz_sources to a static mz_sources to remove an unnecessary function call.
Given we're not sharing iterators across these view generators, we can simplify the code further.
a3430b2 to
b3e9ac9
Compare
mtabebe
left a comment
There was a problem hiding this comment.
I want to make sure I understand this right.
The argument is that we don't need to add the VALUES into the fingerprint of the materialized view because (1) if the schema stays the same we have self correcting behaviour and (2) if the schema does change we may have errors and we will catch them in tests.
The (2) feels fuzzy to me. We will catch them in tests based on our catalog, but what if the customers have different catalogs? Can we be stronger in anyway?
Or are you saying the only errors possible through (2) are schema definition errors which do not depend on the data in the catalog?
Or maybe it is a different argument all together, that we are just moving around definitions?
I guess I'd just like some clarity.
| let materialized_views: &'static BuiltinView = | ||
| Box::leak(Box::new(make_builtin_materialized_views(mv_iter))); | ||
| let tables: &'static BuiltinView = Box::leak(Box::new(make_builtin_tables(table_iter))); | ||
| Box::leak(Box::new(make_builtin_materialized_views(builtin_items))); |
| FROM (VALUES {values}) AS v(oid, schema_name, name, type, privileges)" | ||
| FROM (VALUES {source_values}) AS v(oid, schema_name, name, type, privileges) | ||
| UNION ALL | ||
| SELECT oid, schema_name, name, 'log', privileges |
|
I went through the 7 commits and the code does look good, I'm just trying to make sure my understanding of the error case is correct. I also think this is a change where we should do a nightly run? |
I'd recommend this PR commit by commit. All of this is entirely a refactor and most is just moving code around.
Motivation
Closes sql-653
Description
Should have no customer facing changes. The biggest change is instead of inlining builtin values into the fingerprint, we can use our original pattern of using upstream "mz_builtin_*" views. Initially we thought we had to inline the builtin VALUES into the materialized view, otherwise a silent change to an upstream view (i.e. mz_builtin_indexes) would silently create an error in the materialized view. However we've since realized that as long as the schema of the upstream view stays the same, any updates to it will simply sink into the materialized view. If the schema were to change however, e.g. we ASSERT NOT NULL on the mv but have the upstream view provide a NULL value, then the mv would error when queried. However, we have tests that actually query each builtin so this regression would be caught.
Verification
A proof point of this working is mz_object_dependencies. It is an mv that references an upstream view
mz_object_dependencies_rawthat's constantly changing. This has been in production for a while but things are healthy due to what I stated prior.